Papers with vision-language reasoning tasks
ImageInWords: Unlocking Hyper-Detailed Image Descriptions (2024.emnlp-main)
Copied to clipboard
Roopal Garg, Andrea Burns, Burcu Karagol Ayan, Yonatan Bitton, Ceslee Montgomery, Yasumasa Onoe, Andrew Bunner, Ranjay Krishna, Jason Baldridge, Radu Soricut
| Challenge: | generating accurate hyper-detailed image descriptions is challenging for vision-language models trained on web-scraped image-text. |
| Approach: | They propose a data-centric framework for generating hyper-detailed image descriptions using web-scraped image-text. |
| Outcome: | The proposed framework improves on human evaluations on the data, even with only 9k samples. |
Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Models (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Large Vision-Language Models have demonstrated impressive performance on vision-language reasoning tasks, but their potential for zero-shot fine-grained image classification remains underexplored. |
| Approach: | They propose a method that transforms zero-shot fine-grained image classification into a visual question-answering framework. |
| Outcome: | The proposed method outperforms the current state-of-the-art approach and outperformed existing methods. |